new speaker
Google's mysterious Gemini smart speaker: What we know, and don't know
Blink and you may have missed it, but Google gave us a peek at what sure looks like a new smart speaker during its Made by Google event on Wednesday. A "leaked" product is one that's been mistakenly revealed, whereas the speaker we saw during Google's Pixel event got a clear supporting role, with F1 driver Lando Norris cheerfully chatting with the device. Google meant for us to notice the new and unannounced smart speaker. So, what do we know about this little gray (or porcelain?) That may sound obvious, but so often with rumored or "leaked" new products, we're in the land of pure conjecture.
Google snuck a new smart speaker into its big Pixel event
Over the past several months, the question surrounding Google's next smart devices hasn't been when they will arrive, but if they will arrive. After all, Google's been slowly but steadily discontinuing older smart products (the Nest Protect, the Nest x Yale Lock) while leaving their replacements to third parties. At time same time, its aging line of Nest smart speakers and displays has been languishing. But Google has previously hinted that new Google Home smart devices are on tap for later this year, and during the company's big Made by Google event today, we may have gotten a glimpse of one. During some pre-recorded banter between Milwaukee Bucks player Giannis Antetokounmpo and F1 driver Lando Norris, the camera panned over to reveal a small, slightly squished sphere with a gray exterior and a telltale light ring encircling its narrow base.
Dynamic Recognition of Speakers for Consent Management by Contrastive Embedding Replay
Shahmansoori, Arash, Roedig, Utz
Voice assistants overhear conversations and a consent management mechanism is required. Consent management can be implemented using speaker recognition. Users that do not give consent enrol their voice and all their further recordings are discarded. Building speaker recognition-based consent management is challenging as dynamic registration, removal, and re-registration of speakers must be efficiently handled. This work proposes a consent management system addressing the aforementioned challenges. A contrastive based training is applied to learn the underlying speaker equivariance inductive bias. The contrastive features for buckets of speakers are trained a few steps into each iteration and act as replay buffers. These features are progressively selected using a multi-strided random sampler for classification. Moreover, new methods for dynamic registration using a portion of old utterances, removal, and re-registration of speakers are proposed. The results verify memory efficiency and dynamic capabilities of the proposed methods and outperform the existing approach from the literature. Many recent internet of things (IoT) applications such as smart homes, smart transport systems or smart healthcare rely on voice assistants as primary user interface.
Rapid Connectionist Speaker Adaptation
Witbrock, Michael, Haffner, Patrick
We present SVCnet, a system for modelling speaker variability. Encoder Neural Networks specialized for each speech sound produce low dimensionality models of acoustical variation, and these models are further combined into an overall model of voice variability. A training procedure is described which minimizes the dependence of this model on which sounds have been uttered. Using the trained model (SVCnet) and a brief, unconstrained sample of a new speaker's voice, the system produces a Speaker Voice Code that can be used to adapt a recognition system to the new speaker without retraining. A system which combines SVCnet with an MS-TDNN recognizer is described
Adapter-Based Extension of Multi-Speaker Text-to-Speech Model for New Speakers
Hsieh, Cheng-Ping, Ghosh, Subhankar, Ginsburg, Boris
Fine-tuning is a popular method for adapting text-to-speech (TTS) models to new speakers. However this approach has some challenges. Usually fine-tuning requires several hours of high quality speech per speaker. There is also that fine-tuning will negatively affect the quality of speech synthesis for previously learnt speakers. In this paper we propose an alternative approach for TTS adaptation based on using parameter-efficient adapter modules. In the proposed approach, a few small adapter modules are added to the original network. The original weights are frozen, and only the adapters are fine-tuned on speech for new speaker. The parameter-efficient fine-tuning approach will produce a new model with high level of parameter sharing with original model. Our experiments on LibriTTS, HiFi-TTS and VCTK datasets validate the effectiveness of adapter-based method through objective and subjective metrics.
Call For Speakers - Data, Artificial Intelligence & Advanced Analytics Summit
Speaker Diversity at DPS 2022 Continuing our efforts from last year, the DPS Team will put a considerable amount of effort towards speaker diversity. We will strive hard to achieve the following: Our panels & pool of speakers will represent a gender balance and color diversity % increase in speakers of colors % increase in women speakers Improve gender balance Encouraging & welcoming new speakers assigning mentors additonal support Approach speakers from diverse backgrounds and work towards racial equality Encourage speakers whose first language is not English (DPS 2022 is multi-lingual) Compensate our speakers fairly and equitably Overall, the DPS team is here to support speakers who are women, support BIPOC speakers, engage Asian speakers, expand our reach to the LGTBQIA community, put in extra efforts to reach out to speakers with disabilities (both invisible and visible disabilities). Lastly, encourage speakers whose first language is not English. Yes, DPS 2022 is multi-lingual.
Cross-Lingual Text-to-Speech Using Multi-Task Learning and Speaker Classifier Joint Training
In cross-lingual speech synthesis, the speech in various languages can be synthesized for a monoglot speaker. Normally, only the data of monoglot speakers are available for model training, thus the speaker similarity is relatively low between the synthesized cross-lingual speech and the native language recordings. Based on the multilingual transformer text-to-speech model, this paper studies a multi-task learning framework to improve the cross-lingual speaker similarity. To further improve the speaker similarity, joint training with a speaker classifier is proposed. Here, a scheme similar to parallel scheduled sampling is proposed to train the transformer model efficiently to avoid breaking the parallel training mechanism when introducing joint training. By using multi-task learning and speaker classifier joint training, in subjective and objective evaluations, the cross-lingual speaker similarity can be consistently improved for both the seen and unseen speakers in the training set.
Amazon's 2020 Echo speaker has new features
How does the latest Echo compare to other top smart speakers? The newest thing about the 2020 Echo is its round design. In the box, we find the ball speaker and a power adapter. As with all Echo speakers, you don't need any substantial directions to get it up and running. Just plug it in, open the Alexa app, and follow the prompts when, after a few seconds, the new speaker is detected and a pop-up appears asking whether to set the speaker up.
Everything we know--and don't know--about Google's new smart speaker
The cat's out of the bag as far as Google's latest smart speaker goes, with the company essentially confirming the rumored Home successor on Thursday with a coyly worded email containing a snapshot and a video of the unannounced device. Many key details about the new Google smart speaker are still shrouded in mystery, but we can tease out a few facts by studying the official photo and Google's brief teaser video. Well, no duh, but given that Google gives so little away in its teaser video, we might as well notch this one as one of the few certainties. In the brief video, a man on a sofa says "Hey Google, play some music," and four telltale LEDs on the new speaker (speakers, actually, there two of them) light up as the music begins to play. So yes, Google Assistant confirmed.
Sonos Debuts 3 New Speakers, Including a $799 Soundbar
Nearly two years ago, wireless speaker-maker Sonos released a $400 soundbar that quickly became one of the company's most popular products. The speaker, called Sonos Beam, addressed a few converging trends at the time. It was an easy way to improve the crappy sound coming out of our ever-shrinking flatscreen TVs, it incorporated voice control into the home theater with the inclusion of Alexa and Google Assistant, and its affordable price allowed the company to compete with the flood of cheaper connected speakers on the market. With the Beam established as the best choice for the frugal buyer, Sonos is focusing again on the high end of its product line. The Santa Barbara, California company just revealed the Sonos Arc, a new $799 soundbar that's much sleeker looking than its previous high-end soundbar, the Sonos Playbar.